Systran Mt Dictionary Development

نویسندگان

  • Laurie Gerber
  • Jin Yang
چکیده

SYSTRAN has demonstrated success in the MT field with its long history spanning nearly 30 years. As a general-purpose fully automatic MT system, SYSTRAN employs a transfer approach. Among its several components, large, carefully encoded, high-quality dictionaries are critical to SYSTRAN's translation capability. A total of over 2.4 million words and expressions are now encoded in the dictionaries for twelve source language systems (30 language pairs one per year!). SYSTRAN'S dictionaries, along with its parsers, transfer modules, and generators, have been tested on huge amounts of text, and contain large terminology databases covering various domains and detailed linguistic rules. Using these resources, SYSTRAN MT systems have successfully served practical translation needs for nearly 30 years, and built a reputation in the MT world for their large, mature dictionaries. This paper describes various aspects of SYSTRAN MT dictionary development as an important part of the development and refinement of SYSTRAN MT systems. There are 4 major sections: 1) Role and Importance of Dictionaries in the SYSTRAN Paradigm describes the importance of coverage and depth in the dictionaries; 2) Dictionary Structure discusses the specifics of dictionary structure and types of information represented; 3) Dictionary Creation and Update describes the strategy and mechanics of the dictionary development; 4) Past. Present and Future Development provides some perspective on where SYSTRAN has come from and where it is going. 1. The Role and Importance of Dictionaries in the SYSTRAN Paradigm The problem of automatic linguistic analysis is one of accumulating and utilizing linguistic knowledge. Much of the needed knowledge is static and can be built into the system and its dictionaries, e.g. standard syntax, and conventional word use, including the variations typical of different text types and domains. MT systems encode knowledge in a variety of ways. Linguistic rules, probabilities, and databases of examples are typical encoding methods. SYSTRAN, which is a rule-based transfer system, encodes linguistic knowledge primarily in two ways: 1) In the dictionaries which contain “bottom-up” parse rules along with extensive syntactic and semantic information about each word; 2) In linguistic programs which contain the general “top-down” rules in the parser. Dictionary coverage, or size, is critical to high-quality translation. Words not found in a system's dictionary are a challenge to parse correctly. For SYSTRAN, broad coverage has always been a high priority because of the extensive use of SYSTRAN for information gathering in a wide variety of text types and domains (Gachot 1996.)

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

[Terminology Bulletin 45, 1984, pp.25-37] CHANGES AND IMPROVEMENTS TO THE EUROPEAN COMMISSION'S SYSTRAN MT SYSTEM 1976/84 When the Commission of the European Communities

When the Commission of the European Communities bought its first Systran system in 1976, in what might be called the bronze age of machine translation, it was a system of some 30 000 lines of programming and came with a dictionary of around 6 000 entries. Today, such figures appear laughably small, but for our purposes, Systran was the best there was. Faced with a growing mountain of documents ...

متن کامل

Changes and improvements to the European Commission's Systran MT system 1976/84

When the Commission of the European Communities bought its first Systran system in 1976, in what might be called the bronze age of machine translation, it was a system of some 30 000 lines of programming and came with a dictionary of around 6 000 entries. Today, such figures appear laughably small, but for our purposes, Systran was the best there was. Faced with a growing mountain of documents ...

متن کامل

Comparison of SYSTRAN and Google Translate for English→ Portuguese

Two machine translation (MT) systems, a statistical MT (SMT) system and a hybrid system (rule-based and SMT) were tested in order to compare various MT performances. The source language was English (EN) and the target language Portuguese (PT). The SMT tool gave much fewer errors than the hybrid system. Major problem areas of both systems concerned the transfer of verb systems from source to tar...

متن کامل

Creating a Term Base to Customise an MT System: Reusability of Resources and Tools from the Translator's Point of View

This paper addresses the issue of combining existing tools and resources to customise dictionaries used for machine translation (MT) with a view to providing technical translators with an effective time-saving tool. It is based on the hypothesis that customising MT systems can be achieved using unsophisticated tools, so that the system can produce output of sufficient quality for post-translati...

متن کامل

IWSLT-06: experiments with commercial MT systems and lessons from subjective evaluations

This is a short report of our participation to IWSLT-06. First, we let 2 commercial systems participate as fairly as possible (Systran v5.0 for CE, JE, AE, & IE, Atlas-II for JE), taking care of preprocessing and postprocessing tasks, and tuning as many "pairs" as possible by creating "user dictionaries" and finding a good combination of parameters (such as dictionary priority). Second, we took...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 1997